IndiLem@FIRE-MET-2014 : An Unsupervised Lemmatizer for Indian Languages
نویسندگان
چکیده
An unsupervised and language independent lemmatization procedure has been developed for major Indian languages (Bengali, Hindi etc) which are morphologically very rich and agglutinative in nature. The task of a lemmatizer is mapping an inflected surface word to its appropriate dictionary root word and it is a pre-requisite for implementing several NLP tools like Word Sense Disambiguation system, Machine Translation system, etc. Here the proposed method builts a trie structure using the root words from dictionary and tries to find out a potential lemma of a surface word by efficiently searching in the trie. In this present work, our lemmatization system is tested on Bengali.
منابع مشابه
The IIT Bombay SMT System for ICON 2014 Tools Contest
In this paper, we describe our submission to the ICON 2014 Tools Contest for Machine Translation. The source languages are English, Marathi, Tamil, Telugu, Bengali and the target language is Hindi. We submitted 15 systems; 5 each for the tourism, health and general domains. Our submission is a Phrase-based Statistical Machine Translation system with preprocessing and post-processing elements. A...
متن کاملA Framework for Learning Morphology using Suffix Association Matrix
Unsupervised learning of morphology is used for automatic affix identification, morphological segmentation of words and generating paradigms which give a list of all affixes that can be combined with a list of stems. Various unsupervised approaches are used to segment words into stem and suffix. Most unsupervised methods used to learn morphology assume that suffixes occur frequently in a corpus...
متن کاملA Dictionary- and Corpus-Independent Statistical Lemmatizer for Information Retrieval in Low Resource Languages
We present a dictionaryand corpus-independent statistical lemmatizer StaLe that deals with the out-of-vocabulary (OOV) problem of dictionary-based lemmatization by generating candidate lemmas for any inflected word forms. StaLe can be applied with little effort to languages lacking linguistic resources. We show the performance of StaLe both in lemmatization tasks alone and as a component in an ...
متن کاملOverview of FIRE 2014 Track on Transliterated Search
The Transliterated Search track has been organized for the second year in FIRE. The track has two subtasks. Subtask 1 on language labeling of words in code-mixed text fragments was conducted for 6 Indian languages: Bangla, Gujarati, Hindi, Malayalam, Tamil, Telugu, mixed with English. In Subtask 2 on retrieval of Hindi film lyrics, along with transliterated queries in Roman script, this year we...
متن کاملDPIL@FIRE2016: Overview of the Shared task on Detecting Paraphrases in Indian language
This paper explains the overview of the shared task "Detecting Paraphrases in Indian Languages" (DPIL) conducted at FIRE 2016. Given a pair of sentences in the same language, participants are asked to detect the semantic equivalence between the sentences. The shared task is proposed for four Indian languages namely Tamil, Malayalam, Hindi and Punjabi. The dataset created for the shared task has...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2014